Papers with historical linguistics

5 papers
Automated Cognate Detection as a Supervised Link Prediction Task with Cognate Transformer (2024.eacl-long)

Copied to clipboard

Challenge: Existing methods for cognate identification are based on distributions of phonemes and make little use of cognacy labels.
Approach: They propose a transformer-based architecture inspired by computational biology for automated cognate detection.
Outcome: The proposed architecture performs better than existing methods with increased supervision.
Combining Information-Weighted Sequence Alignment and Sound Correspondence Models for Improved Cognate Detection (C18-1)

Copied to clipboard

Challenge: a new approach to cognate detection is proposed to capture the remaining similarities between cognate word forms after thousands of years of divergence.
Approach: They propose a method which uses information weighting and sound correspondence modeling to improve cognate detection.
Outcome: The proposed approach improves on the measure of form similarity and distance-based cognate clustering.
Cognate Transformer for Automated Phonological Reconstruction and Cognate Reflex Prediction (2023.emnlp-main)

Copied to clipboard

Challenge: Phonological reconstruction is one of the central problems in historical linguistics where a proto-word of an ancestral language is determined from the observed cognate words of daughter languages.
Approach: They propose to use a protein language model to train on multiple sequence alignments to train a model on phonological reconstruction.
Outcome: The proposed model outperforms existing models on cognate reflex prediction task.
Querying a Dozen Corpora and a Thousand Years with Fintan (2022.lrec-1)

Copied to clipboard

Challenge: Large-scale quantitative diachronic corpus studies are difficult if multiple corpus are to be consulted . multi-layer corpus technology can solve the problem, but it requires the user to run queries manually.
Approach: They propose a platform for studying word order in German using syntactically annotated corpora . fintan is a flexible integrated transformation and annotation platform .
Outcome: The proposed platform can be used to study word order in German . it hints at two major phases in the development of scrambling in modern german .
Pater Incertus? There Is a Solution: Automatic Discrimination between Cognates and Borrowings for Romance Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for discriminating between cognates and borrowings are difficult, but they provide a deeper insight into the history of a language and allow for a better characterization of language relatedness.
Approach: They propose a computational approach for discriminating between cognates and borrowings based on a comprehensive database of Romance cognates.
Outcome: The proposed approach is the most comprehensive in terms of covered languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations